Papers by Nurit Cohen Inger

    2 papers
    Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon (2025.emnlp-main)

    Copied to clipboard

    Challenge: Large language models (LLMs) excel on public benchmarks, but high scores may mask overreliance on dataset-specific surface cues rather than true language understanding.
    Approach: They propose a meta-evaluation framework that systematically rephrases benchmark inputs to detect overfitting.
    Outcome: The proposed framework detects performance degradation indicative of superficial pattern reliance on dataset-specific cues and distortion levels.
    DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance (2026.findings-eacl)

    Copied to clipboard

    Challenge: Large Language Models (LLMs) exhibit inconsistent performance across diverse domains.
    Approach: They propose a method that systematically constructs subject-adaptive ensembles by balancing model diversity and competence.
    Outcome: The proposed method achieves 17.1% gain over the best single model, reaching 71.4% accuracy on the MMLU-pro benchmark.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations